61. LLM Optimization
Important optimization techniques include:
- quantization
- batching
- KV caching
- efficient attention
- optimized kernels
- speculative decoding
The goal can be:
lower latency
higher throughput
lower memory
lower cost
Optimization requires measuring the actual bottleneck rather than changing components blindly.